Skip to content

feat(offload): memory.offload_exec=passback — scheduled co-inference of the spilled layers - #809

Open
aldrouil wants to merge 37 commits into
warpfront:masterfrom
aldrouil:feature/cpu-offload-pcie-passback
Open

aldrouil wants to merge 37 commits into
warpfront:masterfrom
aldrouil:feature/cpu-offload-pcie-passback

Conversation

@aldrouil

@aldrouil aldrouil commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Author's Note

This is another relatively large PR and deserves an explanation of its own. This PR directly depends on PR#793 - the partial CPU offload of qwen3.5 dense models.

When I was setting up the offload branch, I noticed that using the CPU to offload compute the layers that were within system RAM was noticeably faster (and it is what llama.cpp does) than just having the GPU access the layers over the PCIe bus. That remains true. But I also noticed that my GPU (a 9070 XT) did not hit 100% utilization while the CPU was computing the layers that were assigned to it. I don't think that this realization is going to surprise most people.

As a result I had a thought and decided to pursue it. The result is this PR, and a feature I am calling 'passback offload' mode.

Passback offload is scheduled co-inference of the layers in system RAM by both the GPU and CPU on a system at the same time. This works and is a speedup if a CPU is unable to saturate the bandwidth between it and the system RAM it is attached to. If there is bandwidth left over, there is sufficient leftover bandwidth for the GPU to read the layers in system RAM over the PCIe bus, and the GPU is fast enough to have some idle time while waiting for the CPU, then this results in an inference speedup of the offloaded layers. In essence, it is a hybrid mode between the existing 'pcie' and 'cpu' offload target modes.

On my machine, which meets all of those requirements, this results in up to a 21.6% speed increase over the 'CPU' offload target pushed in PR#793 depending on model size, and offloaded layer amount. Compared to the 'CPU' offload target, the speedup increases with layers offloaded while the overall speed decreases due to the number of layers offloaded. Better specific benchmarks are available in the PR.

There are a lot of requirements there, so it's worth talking about my personal PC for a moment - because my machine is relatively high end, but it fits a fairly common stereotypical mold.

Specifically, my PC is equipped with:
a) a 7800X3D,
b) 32GB of DDR5 running at 6000MT/s running @ CL30,
c) a B650 Motherboard, which limits my PCIe devices to gen 4.0, and
d) a 9070 XT 16GB.

Those of you who have been paying attention to the computer hardware scene, should note that this is a fairly typical gaming PC setup. It follows the trend of:
a) Fast CPU with less cores than the max available, but also a CPU that has either a high clock speed or in the case of the X3D chips, 3D V-Cache
b) Fairly large but not crazy large amount of decently fast system RAM.
c) A motherboard that may or may not support the latest PCIe generation
d) A GPU that is actually intended only for gaming.

I don't know for certain, but suspect that any PC in a similar "gaming PC shape" to mine will see a similar speedup.

If your PC is running a threadripper or xeon CPU with 96 cores, you probably want to use CPU inference. If you have a dual core CPU running a PCIe 5.0 x16 GPU, you probably want to use 'pcie' offload.

This does add complexity sadly, at the cost of having to have a basic 'scheduler' inside hipfire to allocate the compute of the offloaded layers between the existing 'cpu' and 'pcie' modes in the configuration dynamically. I need to be clear here that this scheduler keeps the layers offloaded to system RAM in system RAM, its only a question of whether the CPU computes inference, or the GPU reads the layer over the PCIe bus.

As before this targets only qwen35 dense models. As before the CPU engine is not byte identical. Further, because the scheduler is dynamic and tries to change the targeted offload amount all the time while running, behavior is not byte deterministic even between runs. By this I mean even though bit parity is not provided between the CPU and GPU runtimes, because passback dynamically allocates layers between the 'cpu' and 'pcie' offload modes output can change from specific run to run - whereas if you offloaded the same number of layers to the cpu runtime you should always get the same output as long as you kept the same number of layers allocated to the cpu runtime.


Summary

memory.offload_exec (env HIPFIRE_OFFLOAD_EXEC) gains a third value, passback, which runs a spilled layer's step on both offload engines concurrently — the GPU arm is pcie's launch restricted to output rows [0, g), the CPU arm is cpu's step restricted to [g, m). The new memory.offload_passback_share (env HIPFIRE_OFFLOAD_PASSBACK_SHARE, auto by default) picks g from the host's measured engine rates and refines it online. The default stays pcie; a fully resident load changes nothing. A step that cannot be split runs wholly on the CPU (the schedule's degenerate point), so the mode is never less robust than cpu.

Behavior, in one line: with memory.offload_exec=passback and a nonzero spill, the GPU stops idling through the host-executed steps and the two engines share each spilled step's output rows.

Which surface(s) does this touch?

  • kernel — crates/hipfire-dispatch (new offload_split module; row-restricted launch views in pipeline/steps.rs; families/gemv.rs row-stride helper), crates/rdna-compute (DType gains PartialOrd/Ord for the per-(dtype, k) schedule map; and, from the bundled fix below, the JIT compile path in compiler.rs — a per-cache-key compile lock)
  • load — crates/hipfire-config (OffloadExec::Passback, PassbackShare + schema rule), hipfire-arch-qwen35 load.rs (CPU-exec coverage line reports splittable layers + the stray-share notice), crates/hipfire-runtime/src/config.rs (retained-Redline exclusion covers passback)
  • serve — crates/hipfire-runtime (llama.rs: the dense FFN down-projection is handed to the generic step seam; new offload_calibrate.rs)
  • arch crate(s): hipfire-arch-qwen35 (load-path coverage report + new tests/gpu_gemv_parity.rs)
  • crates/hipfire-quantize / quant formats — not touched
  • control plane — hipfire-cli (new hipfire offload-bench, --json/--write), hipfire-tui ("Pass-back GPU share" row, exec-mode labels)
  • docs / CI / scripts only — not only docs
  • policy files — not touched

Architecture-trait change? No. crates/hipfire-runtime/src/arch.rs is untouched.

What it is, precisely

  • Both arms are the existing paths, over a row range. The GPU arm is pcie's launch restricted to rows [0, g) (launch_op_rows — a byte view of the weight plus a pointer-offset output; every covered kernel indexes A + row*row_stride and writes y[row]). The CPU arm is cpu's step restricted to [g, m) (cpu_arm_prepare / cpu_arm_finish). No third engine and no new numerics: the GPU arm is bit-identical row-for-row to a full pcie launch, and memory.offload_passback_share=0 is byte-identical to memory.offload_exec=cpu (measured — it is the A/B twin).
  • Ordering, not streams. The blocking D2H is issued before the GPU arm is enqueued (so it drains only the step's producer), the GPU arm runs async, the CPU multiplies while it executes, and the blocking H2D of the CPU's rows is the join — stream-ordered after the GPU arm on the same default stream. No second stream, no events: the copies are k*4 down and (m-g)*4 up (~30 KB at 9B shapes) against tens of MB of weight bytes.
  • The share is scheduled, not hardcoded. The first split-eligible step of each (dtype, k) seeds itself by timing both engines on that step's own weight buffer; every split step refines the share from its own arm timings. The join is read relative to the shape's own measured no-wait floor plus a hysteresis margin, so only the excess over that can be a wait for the GPU arm. Shares are clamped to [0.05, 0.50], adjusted every 4 steps, and a shape latches once converged — but the latch re-opens when the balance point moves (a live reopens counter is on the trace line).
  • Covered shapes: Gemv{Raw}, Gemv{Prerotated}, and either residual form when the dtype has a fused residual kernel. Dense qwen3.5 only; MoE/routed paths never reach it. Everything else runs on the CPU engine alone.
  • Eligibility gates (a failure is the schedule's degenerate point, never pcie): not host-mapped; padded row_stride; weight byte length not exactly m * row_bytes(q, k); no vector row dot for the format; non-F32/short output; weight below 2 MiB; residual form with no fused GPU residual kernel.
  • Accounting stays honest. A split step is not charged to the CPU-idle numerator and gets its own split: … line under HIPFIRE_CPU_EXEC_TRACE=1, which also prints the controller state (gpu samples, floor, waited, target, applied, reopens, frozen).
  • New hipfire offload-bench measures the host directly (one row per dense format: cpu GB/s, gpu GB/s, the share, the two-engine speedup bound), with --json and --write to persist the recommendation. It is a raw device benchmark and refuses while a daemon pid file names a live process. Its measured limits — the shape matrix and the --write pin — are in "Exposed by this work" below.
  • Design doc: docs/plans/partial-gpu-offload-design.md § 6.2.2. Config keys in docs/CONFIG.md / docs/env-vars.md (regenerated by scripts/check-lifecycle.py --write); changelog under Unreleased.

Measured (gfx1201, RX 9070 XT, PCIe 4.0 ×16, Ryzen 7 7800X3D)

Read this as fixture-bound measurement, not a product default. Absolutes are
contended-host; only the interleaved relative deltas transfer. Protocol and
caveats: docs/methodology/perf-benchmarking.md;
each number lives in an immutable historical checkpoint record (list below).

model / quant spilled cpu passback pcie passback vs cpu paired rounds
2B mq4 (24 L) 6/24 69.9 79.3 101.0 +13.4 % 5/5 +
2B mq4 12/24 42.1 47.1 60.2 +11.9 % 5/5 +
2B mq4 18/24 30.0 34.0 42.6 +13.3 % 5/5 +
9B mq4 (32 L) 8/32 28.1 33.6 23.8 +19.6 % 5/5 +
9B mq4 16/32 16.5 19.9 13.2 +20.6 % 5/5 +
9B mq4 24/32 11.6 14.1 9.1 +21.6 % 5/5 +
9B mq3 (32 L) 8/32 27.0 31.1 27.8 +15.2 % 5/5 +
9B mq3 16/32 16.5 19.7 16.2 +19.4 % 5/5 +
27B mq4 (64 L) 4/64 20.6 22.5 17.8 +9.0 % mixed (see record)
27B mq4 8/64 15.0 16.9 11.7 +12.7 % 3/3 +
27B mq4 16/64 9.6 11.3 6.9 +17.7 % 3/3 +

Data: docs/perf-checkpoints/2026-10-02-offload-passback-model-spread-gfx1201.md (+ amendment-1),
…-quant-spread-gfx1201.md (+ amendment-1),
…-mode-ab-gfx1201.md,
…-split-gfx1201.md,
…-headroom-idle.md.

What the numbers say:

  • pass-back beats cpu on every model and quant measured (+9 % to +22 %), and every paired round was positive — so over cpu the gain is roughly flat across size and quant. What changes is which single-engine route it beats: pcie is the loser on the 9B/27B but the winner on the 2B, so the recommended memory.offload_exec is size- and quant-dependent.
  • Co-inference is near its floor: the largest 9B step runs within ~5–10 % of the two-engine balanced bound; d2h+join is ≈ 21 % of the step wall. Refused (never-split) steps are only 3.7–6.3 % of CPU-executed wall — coverage is not the deficit (…-coverage-phase0.md + amendments 1–3).
  • Headroom this exists to use: the GPU is idle 77.7 % of a cpu step's wall at 8/32 spilled, 89.8 % at 16/32, 94.6 % at 24/32, 66.0 % on a 27B at 8/64; a CPU GEMV stream and a GPU PCIe read of the same host-mapped bytes are near-additive to ~75–78 GB/s.
  • Dense FFN down-projection (fused SiLU + down GEMV + residual, the one op family the seam could not see) now hands its residual GEMV to the seam as a Step::GemvResidual: measured +3.8 % at 8/32 and +2.6 % at 16/32 (9B, interleaved fresh-process pairs).
  • Latch re-open tuning is a disclosed null. The four-arm re-open band A/B came back inside its own run-to-run spread and flipped the ordering between budgets, so no perf win is claimed for it; it rests on the freeze-band margin and the drift model (…-latch-band-ab.md + amendment-1). The mode's measured headline does not depend on it.

Figures (committed):

Determinism (please read before diffing output)

The modes are not byte-interchangeable, and passback is the only one that is
not run-to-run reproducible:

  • cpu and pcie run a layer's GEMV on different engines, so they need not agree — expected, not a bug (measured on the 2B: cpu 2046 B vs pcie 3554 B). Within a mode at the same layer split the output is deterministic (cpu×2 and pcie×2 byte-identical).
  • passback schedules its row split per process (seeding probe + online arm timings), so two identical greedy invocations mix the engines differently. Pin memory.offload_passback_share in any gate that diffs pass-back output against a reference, or use 0 to reach the exact cpu twin.
  • Consequence documented in AGENTS.md's pitfall table: never assert a greedy/golden reference across a mode pair.

Exposed by this work, not fixed here

  • qwen3.5:9b-mq6 does not decode in a spill configuration, for a reason unrelated to pass-back: its tensors are legacy MQ6G256, whose plan is the unfused [SiluMulRotate, GemvResidual], but weight_gemv_swiglu_residual takes the fused launcher and dispatch_swiglu_residual has no MQ6G256 arm — the catch-all fires and every mode (passback/cpu/fully resident) fails with unsupported gemv.swiglu_residual for /. Flagging it for a separate fix; the pass-back PR does not touch it.
  • Scheduler estimator limitations (documented in the design doc, not defects of this PR, and the next work in this path): min_join_ns is a monotone running minimum that never rises (a contention-lifted floor biases the balance point down); r_gpu is refreshed only on a waited join, so a shape whose joins sit at the floor carries a stale GPU-rate anchor. Neither is where this fixture's gap is: the 2026-10-01 investigation measured the share sitting where its probe said it should in that regime (docs/investigations/2026-10-01-offload-passback-perf-gap.md § 5), and the skipped-step wall term is 3.7–6.3 %. (That regime's probe agreement does not carry to the default offload-bench matrix on today's host — see the next bullet.)
  • hipfire offload-bench's default shape matrix does not represent a real projection, and --write is a hard pin. The probe times each engine alone on a synthetic Step::Gemv{Raw} shape — k=5120, m derived from --buffer-mb (192 MiB → m≈37k–97k rows), never a real projection and never Step::GemvResidual — so its recommendation is a solo-regime bandwidth ratio. Measured on gfx1201, 2026-10-02: it recommends 0.331 for mq4g256 (mq4g256v2 0.334, mq4cg256 0.331, mq6g256 0.395, mq5g256v2 0.450, q8_0 0.313), while the scheduler's own per-shape values on the same host are 0.489–0.492 for the non-residual shapes and 0.353 for the FFN down-projection; --k 4096 --buffer-mb 4 reproduces the live number (0.494). Two further notes on that run: the recommendation is taken over the five formats that win a ~1.6 KB tie among equal 192 MiB measurements (median over all nine: 0.345), and the per-format number moves ~±0.05 run to run. Because --write pins the value for every shape (offload_split.rs: Share(f) => f skips the Auto branch where the seeding probe and the controller run; the pin is only recorded for the trace) — unlike seed(), which installs a fallback the controller still refines — writing that default here would hold the non-residual steps ~0.16 low. auto is the right setting on this host; --write is not. A realistic default shape, a residual probe mode, and a --write warning are separate follow-up work, not part of this PR.

Test plan

  • ./scripts/no-gpu-ci.sh passes — after the bundled JIT fix below; the post-fix run was end-to-end with HIPFIRE_KERNEL_CACHE pointed at a fresh directory (exit 0, 285 s). Before the fix, a cold-cache run failed 5 rdna-compute GPU tests on a clang-22 frontend crash (details and fix in "Bundled CI fix"); everything else in the script passed on the first run (the script exits at the first failing step, so those steps were run individually and are green: hipfire-cpu 26, hipfire-arch-qwen35 moe_prefill 20, config/registry/client/cli/tui 165, pytest 388+6 subtests, redline unittest 179, test_install_revision.py, test_uninstall.py, check-env-docs.py, check-lifecycle.py 268 keys / 1356 env vars).
  • cargo check --workspace --examples clean (first step of no-gpu-ci.sh); cargo build --release run to completion (47 s incremental, warnings only, no errors).
  • cargo test --lib --workspace as a whole was not run — no-gpu-ci.sh runs the no-GPU subset (rdna-compute, hipfire-cpu, hipfire-arch-qwen35 moe_prefill, config/registry/client/cli/tui) and that subset is green. Flagging the gap rather than claiming the wider command.
  • load / serve / kernel changes: I ran scripts/serve_harness.py on hardware myself for the exact model and settings under test, and attached the per-turn JSON below (passback battery + chain, plus cpu and fully-resident controls on the same fixture).
  • KV backend fields inspected: every attached turn reports kv_backend: "vmm", kv_backend_legacy: false, kv_backend_reason: null. No legacy token was used.
  • serve_harness is ticked as user-facing serve semantics only, per docs/VALIDATION.md's claim→route map — it is explicitly not numerical/state parity evidence, and no cross-mode golden-output assertion is possible on this path: cpu and pcie run different engines, and passback schedules its row split per process, so two identical invocations mix the engines differently. What the offload path's parity rests on instead: (a) memory.offload_passback_share=0 is byte-identical to cpu (measured, and the A/B twin), (b) the GPU arm is bit-identical row-for-row to a full pcie launch by construction — the same kernels over a row range, no new numerics (the design doc's claim; unlike (a) it is not a separate measured byte-diff), and (c) crates/hipfire-arch-qwen35/tests/gpu_gemv_parity.rs — the GPU↔CPU GEMV parity matrix for the host-mapped path, new in this branch, #[ignore]d because it needs a GPU.
  • If perf-relevant: ./scripts/speed-gate.sh — not run. Its locked baselines (tests/speed-baselines/gfx1201.txt) were captured on a 4× R9700 / Threadripper host; this host is a single RX 9070 XT, so a comparison would report the hardware, not the change. The perf claim here is fixture-bound and lives in the checkpoints above; the resident control in the harness evidence is the regression smoke for the shared step seam.
  • No ceiling in scripts/leanup-thresholds.txt is raised by this branch (the file is untouched), so no ratchet-raise label is owed.
local serve_harness harness output (load / serve / kernel changes)

Fixture: ~/.hipfire/models/qwen3.5-9b.mq4 (md5 296092bf1e6a45d78c1acf815eb93366), greedy,
--speculation off --thinking off, spill via HIPFIRE_GPU_LAYER_BUDGET=24
(24 resident of 32 → 8 spilled); the passback/cpu arms add
HIPFIRE_OFFLOAD_EXEC, the resident control runs the same model with the budget
unset. Daemon md5 d8e21f57834240b7607dba80546be863; hipfire md5
31c057d13c637446aafab0685e8035e6 — the build before the bundled JIT fix
below (post-fix build: daemon 80e52b2c3655a55e051f783563491729, hipfire
7bc03ee5898c28abfd13c43dde42f551). The fix is behavior-neutral for this
evidence: it only serializes duplicate compiles of one cache key and touches no
forward-path code.

--mode battery, HIPFIRE_OFFLOAD_EXEC=passback:

[
 {
  "request_id": "chatcmpl-846787-1",
  "finish": "stop",
  "gen": 343,
  "decode_tok_s": 33.2,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "retrieval_missing": [],
  "expected_substrings": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_reason": null,
  "prompt_md5": "43ca0d15712d3dfb777b51ae76d8fd5f",
  "request_md5": "b2d411eb58ec2648ebd087ac828a51bb",
  "ans_preview": "```python\ndef merge_sorted(a, b):\n    \"\"\"\n    Merge two already-sorted lists into one sort"
 },
 {
  "request_id": "chatcmpl-846787-3",
  "finish": "stop",
  "gen": 286,
  "decode_tok_s": 33.6,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "retrieval_missing": [],
  "expected_substrings": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_reason": null,
  "prompt_md5": "640e0fd4f55996cb175a422f0a12cef5",
  "request_md5": "85b33e9aedabf7a9fa0d1d81cc6f8356",
  "ans_preview": "To find the total distance traveled, we need to calculate the distance for each segment of"
 },
 {
  "request_id": "chatcmpl-846787-5",
  "finish": "stop",
  "gen": 81,
  "decode_tok_s": 33.5,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "retrieval_missing": [],
  "expected_substrings": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_reason": null,
  "prompt_md5": "8f66b4c97988825bd8e7840aaf44357e",
  "request_md5": "d34383b66fd5f7af0cfced19d3777983",
  "ans_preview": "The primary cause of Earth's seasons is the tilt of its rotational axis relative to its or"
 },
 {
  "request_id": "chatcmpl-846787-7",
  "finish": "stop",
  "gen": 102,
  "decode_tok_s": 33.6,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "retrieval_missing": [],
  "expected_substrings": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_reason": null,
  "prompt_md5": "8fe0ad36f61bcf4992cc9df81cdf3817",
  "request_md5": "643a78544be2dab1edcfe5b8e85c3718",
  "ans_preview": "Elias tended his lonely lighthouse for decades, watching the endless waves crash against t"
 },
 {
  "request_id": "chatcmpl-846787-9",
  "finish": "stop",
  "gen": 82,
  "decode_tok_s": 33.5,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "retrieval_missing": [],
  "expected_substrings": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_reason": null,
  "prompt_md5": "8bed8e2d056dc1d47dccae9d32dbecf4",
  "request_md5": "9af4a6122e824b100c1d31cb525d7816",
  "ans_preview": "1. Use meaningful variable and function names that clearly describe their purpose.\n2. Keep"
 }
]

The cpu and fully-resident controls run the same battery prompts as the
passback battery above, turn for turn — their request_md5 values match it
exactly, which is what makes the three battery arms comparable. chain is the
related-turn mode, so from turn 2 onward it carries prior turns into the request
and its request_md5 values legitimately differ (efde613f… at turn 2 vs the
battery's 85b33e9a…).

arm turn finish gen decode tok/s attractor empty runaway retrieval_missing kv_backend
passback chain 1..5 stop 343 / 267 / 85 / 90 / 90 33.3 / 32.8 / 33.1 / 32.9 / 33.3 false false false [] vmm, legacy false
cpu battery 1..5 stop 343 / 286 / 81 / 102 / 82 27.9 / 27.9 / 28.0 / 28.1 / 28.0 false false false [] vmm, legacy false
resident battery 1..5 stop 343 / 286 / 81 / 102 / 82 120.7 / 121.3 / 122.1 / 121.7 / 121.7 false false false [] vmm, legacy false

Decoded text read: coherent on all four arms. On this fixture passback
reproduced the cpu and fully-resident text exactly — all five battery
turns byte-equal (assistant_content diffed turn-by-turn, 5/5 against each arm),
and the same token counts (343 / 286 / 81 / 102 / 82). That is a property of this
fixture's argmax, not a guarantee: passback's row split is scheduled per
process, which is why the determinism section above forbids building a
golden-output gate on it. passback 33.5 vs cpu 28.0 tok/s here is +19.6 % —
the same delta as the 8/32 row of the checkpoint table. Full per-turn
JSON (every field) for all four arms was kept at /tmp/passback-pr/{pbn_battery, pbn_chain,cpuN_battery,residentN_battery}.json on the author's host.

Bundled CI fix (independent of pass-back) — the reviewer may prefer this split out

rdna-compute's JIT path launched one hipcc per caller for the same
module whenever several threads needed it on a cold kernel cache (the
rdna-compute lib test binary runs its GPU tests on a thread pool, and every
test builds its own Gpu). Every failing run had exactly that condition, and
the ROCm 7.2 clang-22 frontend crashed in it — a parser stack fault,
clang::Parser::ParsePragmaLoopHint → SIGSEGV in
clang::Lexer::LexTokenInternal, current parser token 'xa' at
tensor_ops.hip:1332 inside hyper_norm_gate_f32. The source file on disk is
intact and byte-identical between the failing and the passing cache
(md5 37119ab9…, 160694 B), so this is not a torn read of the shared
{stem}.hip; it is the compiler crashing, intermittently, while identical
compiles ran.

compiler.rs now takes a process-global lock per cache key across the
compile (compile_lock), and a waiter re-checks the published blob and reuses
it instead of compiling the same source again. Distinct keys still compile in
parallel — that is compile_batch_for_symbols's contract — so this is not a
global compile lock. It removes the redundant concurrent compiles (5 identical
hipcc runs become 1) that were present in every failing run.
The compiler ICE itself is not claimed fixed — it is intermittent, was not
reproduced by hand, and the backtrace below is worth reporting to LLVM/ROCm.

Evidence — cargo test -p rdna-compute --lib (400 tests, HIPFIRE_KERNEL_CACHE
pointed at a fresh directory so every module JIT-compiles):

run result
before, warm-ish cache (scripts/no-gpu-ci.sh, the reported failure) 5 failed / 393 passed
before, cold cache 1 failed / 397 passed
after, cold cache ×3 0 failed / 399 passed, three times (254 s / 259 s / 254 s)
after, scripts/no-gpu-ci.sh end to end, cold cache exit 0 (285 s)
after, cargo test -p rdna-compute --lib compiler:: 35 passed (incl. the new compile_lock_is_per_cache_key_and_mutually_exclusive)

Three clean runs do not prove an intermittent compiler ICE is gone — they are
the strongest evidence available on one host. What is proven is the
deterministic half: one compile per cache key, and the new unit test asserts the
lock is per-key and mutually exclusive.

clang-22 crash backtrace (first frames)
1. /tmp/kc-cold-2/gfx1201/tensor_ops.a0c5e4df29aefe9b.hip:1332:26: current parser token 'xa'
2. /tmp/kc-cold-2/gfx1201/tensor_ops.a0c5e4df29aefe9b.hip:1249:21: parsing function body 'hyper_norm_gate_f32'
3. ...:1249:21: in compound statement ('{}')
4. ...:1324:65: in compound statement ('{}')
5. ...:1329:41: in compound statement ('{}')
 #0 llvm::sys::PrintStackTrace(llvm::raw_ostream&, int)
 #1 ...
 #2 (/usr/lib/libc.so.6+0x44cb0)
 #3 clang::Lexer::LexTokenInternal(clang::Token&, bool)
 #4 clang::Preprocessor::Lex(clang::Token&)
 #5 clang::Preprocessor::PeekAhead(unsigned int)
 #6 clang::Parser::ParseDeclaratorInternal(clang::Declarator&, void (clang::Parser::*)(clang::Declarator&))
 #7 clang::StackExhaustionHandler::runWithSufficientStackSpace(...)
 ...
#22 clang::Parser::ParsePragmaLoopHint(llvm::SmallVector<clang::Stmt*, 24u>&, clang::Parser::ParsedStmtContext, clang::SourceLocation*, clang::ParsedAttributes&, clang::LabelDecl*)

Hardware validation request (optional)

Deliberately omitted. The claim needs HIPFIRE_OFFLOAD_EXEC=passback plus a
memory.gpu_layer_budget spill, and the <!-- hw-gate-request --> block's
routes[] schema cannot express env for a route — a hw-gate run on the default
fixtures would exercise pcie, not this mode, and report a non-reproduction that
is not a result. Local serve_harness output is attached instead. If a
maintainer wants it on automation, the route is: qwen3.5:9b battery + chain
with HIPFIRE_OFFLOAD_EXEC=passback HIPFIRE_GPU_LAYER_BUDGET=24 (fixture:
~/.hipfire/models/qwen3.5-9b.mq4, md5 296092bf1e6a45d78c1acf815eb93366).

How this merges (direct review)

Merge authority is direct maintainer review plus the required no-GPU CI
checks. hw-gate automation is optional evidence delivery, not a prerequisite or
a substitute.

  1. Required CI (must stay green): build (workspace, no GPU), unit tests (lib, no GPU), gates (ratchets, layering, registers) from .github/workflows/ci.yml.
  2. Required review: one approving maintainer review, judging claim-matched evidence for the surfaces ticked above (static read for docs; hardware/model proof for the load/serve/kernel behavior).
  3. Evidence owed: routes from docs/VALIDATION.md; the local serve_harness output and the perf checkpoints are attached/linked above.
  4. Optional automation: hw-gate seat commentary is not a required status check and does not auto-approve or promote master.
  5. Retired paths: scripts/coherence-gate*.sh and tools/change_gate / agentic-review are historical only — never acceptance evidence.

Architecture-trait change?

No. crates/hipfire-runtime/src/arch.rs is untouched; no Architecture trait
surface changes, so nothing ripples to the arch crates beyond
hipfire-arch-qwen35's own load-path coverage line.

Avery Drouillard added 30 commits October 1, 2026 10:00
…om record

`HIPFIRE_CPU_EXEC_TRACE=1` now also prints, on the global step count's doubling
schedule, `cpu exec: idle N% — window ending at step S ...`: the share of a decode
window's wall spent inside CPU-executed steps. A CPU step is a host sync point
(the blocking D2H, the multiply, the H2D) and prefill never enters the seam, so
that share is a lower bound on the GPU's idle fraction — the headroom a scheme
that hands part of a spilled step back to an idle GPU would be spending.
Windowed rather than cumulative, because `hipfire bench --runs N` decodes N times
inside one process and the gaps between runs are neither.

Measured with it (gfx1201, RX 9070 XT, PCIe 5.0 x16, HIP 7.2, daemon md5
7f115e384f4fe1521d02dc57fbd02745): the idle share rises with the CPU's share of
the split — 9B 77.3/77.7/77.8% at 8 of 32 spilled, 89.8% at 16, 92.4% at 20,
94.6% at 24; 27B-mq3-xt 66.0% (repeat 66.3%) at 8 of 64. A four-process
contention probe (CPU gemv stream over a 205 MB host-mapped buffer, PCIe H2D of
the same bytes, then both) shows the two streams are close to additive: the CPU
retains 86-94% of its solo rate while the GPU runs at 102-103% of its, combined
75-78 GB/s — ~1.4x CPU-alone and ~90% of this host's DDR5 peak. So the idle
window is real, and a row-split pass-back would be bounded by host DRAM rather
than by either engine.

Record and raw captures: docs/perf-checkpoints/2026-10-01-offload-passback-headroom-idle.md
and its data directory (bench JSON, per-window idle lines, the probe source).
Rates there are n=1 per process and labelled exploratory; the idle fractions and
the probe are the evidence. The record also documents the 27B load hang (layer
45/64, 99% CPU, 0 GB free) as the capacity precondition the 2026-09-27 amendment
already flagged, and that the 2026-09-27 9B fixture is not in any registry entry
or HF revision, so both arms are measured fresh on the registry artifact
(qwen3.5:9b, sha256 ba83acf5...).

Diagnostic only: no behavior change to the resident or offload paths, no new env
var. Docs updated in env-vars.md, partial-gpu-offload-design.md 6.2.1 and the
benchmark handoff.
LLM incorrectly determined that GPU was a PCIe 5.0 x16 link. This is
correct in what the GPU is capable of, but incorrect in what it actually
has available. Both the GPU and CPU on this system support a PCIe 5.0
x16 link - the B650 motherboard does not. As a result the link is capped
at PCIe 4.0.
…ross both engines

A CPU-executed spill step is a host sync point (blocking D2H, host GEMV,
blocking H2D), so the GPU is idle for most of its wall: 77.7 % at 8 of 32 layers
spilled on the 9B, 94.6 % at 24 of 32, 66.0 % on a 27B at 8 of 64
(docs/perf-checkpoints/2026-10-01-offload-passback-headroom-idle.md). The same
record measured the headroom: a CPU GEMV stream and a GPU PCIe read of the *same*
host-mapped bytes are near-additive to 75–78 GB/s combined against ~46 GB/s for
the CPU while the GPU reads.

A spilled step's output rows are independent, so this mode splits the step by
output rows: rows [0, g) go back to the GPU (the production `pcie` kernel on the
same host-mapped weight) and rows [g, m) stay on the CPU, concurrently.

* Both arms are the *existing* paths over a row range, so there is no third
  engine and no new numerics — and `memory.offload_passback_share = 0` is
  byte-identical to `memory.offload_exec=cpu`.
  - `launch_op_rows(gpu, ctx, step, Some(0..g))` restricts a `pcie` launch: a byte
    view of the weight's rows plus `sub_offset` views of the output and residual.
    Every covered kernel indexes `A + row*row_stride` from the passed base and
    writes `y[row]`, so a row-shifted base *is* a row-shifted launch. `None` is the
    pre-existing path, unchanged and still the only one `pcie` uses.
  - `cpu_exec::cpu_arm_prepare` / `cpu_arm_finish` are `cpu` mode's step in its two
    observable phases; `run_host_mapped_gemv{,_residual}` and `run_step` keep their
    signatures and run the two back to back, so `cpu` mode's device-operation
    order is unchanged.
* Ordering, not streams. The blocking D2H is issued before the GPU arm is
  enqueued (so it drains only the step's producer), the GPU arm runs async, the
  CPU multiplies while it executes, and the blocking H2D of the CPU's rows is the
  join — stream-ordered after the GPU arm on the same stream. The copies are
  `k*4` bytes down and `(m-g)*4` up against tens of MB of weight bytes, so a
  second stream plus events would buy the overlap of a ~2 µs copy.
* `memory.offload_passback_share` (`HIPFIRE_OFFLOAD_PASSBACK_SHARE`): `auto`
  (default) schedules the share, a number in (0, 0.5] pins it, `0` disables the
  pass-back. Inert outside `passback` (one informational line at load). Requires
  the same config plumbing as the mode: `OffloadExec::Passback`,
  `ValueRule::PassbackShare`, the TUI row/knob and the schema arms.
* The share is scheduled, not hardcoded — the optimum is a property of the host,
  so nothing upstream depends on this box's constant. The first split-eligible
  step of each (dtype, k) seeds itself by timing both engines on that step's own
  weight buffer (GPU arm through a scratch-output probe, so a residual probe
  cannot be double-counted), and every split step refines the share from its own
  arm timings: `r_cpu` from the contended CPU arm every step, and `r_gpu` only
  when the join outlasts the CPU's own multiply — the one regime where the
  blocking H2D demonstrably waited for the GPU arm. Below it the join *is* the
  CPU's duration, so that sample would read ~20× low and pin the split to its own
  floor with both engines idle in turn; there the share rises instead. Clamped to
  [0.05, 0.50], adjusted every 4 steps, frozen once a shape has applied 64
  adjustments with the last four under 0.005.
* A failed seeding probe degrades to DEFAULT_GPU_SHARE rather than failing a step
  that plain `cpu` mode would have run, and every refusal falls back to the
  whole-CPU step — never to `pcie`. Refused: a padded `row_stride`; a weight whose
  byte length is not exactly `m * row_bytes(q, k)` (the invariant that makes
  `&host_bytes[g*row_bytes..]` the CPU arm's row `g`); no vector row dot (with the
  scalar decoder the best share is all-GPU, which `pcie` does better); a non-F32 or
  short output; a weight below 2 MiB; and a residual step with no GPU residual
  kernel (`dispatch_residual`'s dtype set, which is `for_gemv_residual`'s).
* Accounting stays honest: a split step is not charged to the CPU-idle numerator
  (its wall contains GPU work), gets its own `split:` trace line under
  `HIPFIRE_CPU_EXEC_TRACE=1` (`join` = the blocking H2D, `cpu=`/`gpu≥` the sampled
  and bounded rates), and the `cpu exec: idle` window names the split steps it
  excluded.
* `hipfire offload-bench` measures the host directly (one row per dense format:
  cpu/gpu GB/s, the share, the two-engine speedup bound; `--json`, `--write`),
  building its own buffers via `offload_split::probe_synthetic` and refusing while
  a daemon pid file names a live process. It lives in `hipfire-runtime` so the CLI
  grows no dependency.

Evidence: `cargo test -p hipfire-config` 101, `-p hipfire-dispatch` 310,
`-p hipfire-tui` 165 pass. New `#[ignore]`d row-offset equivalence arms
(`tests/gpu_gemv_parity.rs`) are bit-exact on gfx1201 over 15 cases — real 2B/9B
MQ4/MQ3-Lloyd/MQ6/HFQ6 tensors, their odd-row-count prefixes, and synthetic
MQ4G256V2/MQ4CG256/MQ3G256V2/HFQ4G256/Q8_0 — for both arms of both step forms,
including that rows past `m` are never written and that a residual shape with no
GPU arm is refused. `memory.offload_passback_share=0` output is byte-identical to
`memory.offload_exec=cpu`, and `pcie` output is byte-identical with and without a
stray share set.
…I row, AGENTS, changelog

* `docs/plans/partial-gpu-offload-design.md` § 6.2.2: the row split, why ordering
  rather than a second stream buys the overlap, the four covered shapes, the
  refusal list, the scheduler (probe + the join-vs-gemv online signal), the share
  key, and the measured ceiling from the 2026-10-01 headroom record.
* `docs/CONFIG.md` / `docs/env-vars.md`: the generated lifecycle/inventory rows
  (`scripts/check-lifecycle.py --write`) plus the hand-written prose row for
  `HIPFIRE_OFFLOAD_EXEC` (`passback`) and a new
  `HIPFIRE_OFFLOAD_PASSBACK_SHARE` row.
* The pass-back share is described for operators as a scheduled value with a
  pinned override, and as the byte-identity A/B twin at 0.
* `AGENTS.md` § 7 flag table gains `HIPFIRE_OFFLOAD_PASSBACK_SHARE` and names the
  third `HIPFIRE_OFFLOAD_EXEC` value.
* `CHANGELOG.md`: an `## Unreleased` entry for the mode.

Gates: `scripts/check-lifecycle.py` and `scripts/check-env-docs.py` both exit 0.
… join against its own floor

Two measurement bugs found while verifying `memory.offload_exec=passback`, both of
which produced plausible-looking wrong numbers rather than errors.

1. **`sync_with_deadline` cannot time anything.** It polls completion with
   `std::thread::sleep(SYNC_POLL_INTERVAL)`, `SYNC_POLL_INTERVAL = 2 ms`
   (`crates/rdna-compute/src/dispatch.rs:1264`), so a probe that times
   `launch → sync_with_deadline` measures the *sleep*. The seeding probe and
   `hipfire offload-bench` therefore reported 4.1 GB/s for an 8 MiB host-mapped
   weight whose true rate is 24.4 GB/s, and the artifact looked exactly like a
   size-dependent fixed cost (~1.7 ms/launch) — which is how it survived a first
   read: `--buffer-mb` 8/16/24/48/96/192 gave 4.1/8.1/12.2/12.2/16.3/24.4 GB/s,
   a textbook `a + b·bytes` fit, and the fit's slope (~32 GB/s) was right while
   its intercept was the poll interval. `offload_split`'s probe now uses
   `gpu.hip.device_synchronize()`; deadline-bearing paths keep the polling sync.
   The seeded share moved from ~0.23 to ~0.36, and the bench's GPU column became
   size-independent (24.4–26.6 GB/s over 8 MiB–192 MiB).

2. **The controller read the join as a GPU duration in the regime where it is the
   CPU's.** The join is `copy + max(0, gpu_ns − cpu_ns)`, sampled as
   `rows_gpu·row_bytes / (cpu_ns + join_ns)`. Whenever the CPU arm is the straggler
   (the common case: a 0.1–0.6 ms CPU multiply against a 30–50 µs copy) that
   denominator is the *CPU's* duration, so the sample read ~13× low — and since a
   seeded `r_gpu` then wins `next_share`'s both-rates arm forever (the EWMA only
   feeds on waits), it pinned every shape to 0.071–0.081, just above `SHARE_MIN`,
   with both engines idle in turn. Any absolute microsecond threshold has the same
   failure on this host, because the copy alone costs more than the 20 µs the
   threshold assumed.
   The join is now measured against the shape's **own** no-wait floor (the smallest
   join seen) plus a hysteresis margin (10 µs, or a quarter of the floor): only the
   excess over that can be a wait, `gpu_ns = cpu_ns + excess`, and at the floor the
   share rises instead — the direction the design always intended. Measured shares:
   0.279–0.360 on the 9B's field shapes, against 0.28–0.45 by format and 0.36–0.40
   for the same sizes from the standalone bench.

Also in this commit:

* The `split:` trace line now prints the controller's own state (`gpu samples`, the
  last `waited`, the last `target`, `applied`, `frozen`). A share nobody can explain
  is diagnosable from one run; the two bugs above took one run each to localise
  because of it.
* `offload-bench`'s recommendation is the **median** share over the largest
  measurements. Every format is measured over the same `buffer_mb`, so "the largest
  measured format" is normally the whole matrix and `max_by_key(bytes)` degenerated
  to whichever format happened to be last.
* The residual gate refuses *any* `GemvResidual` whose dtype has no GPU residual
  kernel, not just the `Raw` form: the `Prerotated` form would have failed to launch
  where `cpu` mode handles it. (Found by the parity test's Q8_0 case.)
* The row-offset parity arms report the **two arms' mutual delta** and require at
  least one case where the engines' numbers differ, so "the split mixes both
  engines" is asserted rather than assumed: 12 of 15 cases differ (1.2e-7–1.5e-5).

Verification (gfx1201, RX 9070 XT, 9B MQ4 registry artifact, budget 24 = 8 of 32
spilled, greedy, `--spec off`, pinned prompt md5 5835c71e…):
* 15 parity cases bit-exact per arm, plain and residual forms, including no writes at
  or past `m`.
* `share = 0` output byte-identical to `cpu` (555-byte and 642-byte completions,
  `cmp`), and no `split:` line at all.
* `pcie` with a stray share byte-identical to `pcie` alone, one notice line.
* `hipfire bench`, three interleaved fresh-process rounds: `cpu` 27.40, `passback`
  30.60 (+11.7 %), `pcie` 23.20 tok/s; `vram_free_mb` identical across arms.
* `offload-bench` exit 0 with sane rates (GPU 23.0–28.2 GB/s vs the record's 27.3;
  CPU 34.7–68.4 vs 52), and exit 1 naming the live pid while a daemon runs.
* `-t 0 --spec off -n 256/512` text: `cpu`, `passback` and `pcie` all byte-identical
  on this fixture (the 2026-09-27 cpu-vs-pcie divergence was measured on a different,
  local-only 9B).
…8 of 32 spilled)

Dated, fixture-bound `historical` record for `memory.offload_exec=passback`, the
companion of the headroom record that motivated the mode. One host (RX 9070 XT,
PCIe 4.0 x16, Ryzen 7 7800X3D, HIP 7.2), one model (registry 9B MQ4, sha256
ba83acf5…), one prompt (md5 5835c71e…), one budget (24 of 32 layers resident),
greedy `--spec off`.

Establishes: +11.7 % decode over `cpu` on the median of three interleaved
fresh-process rounds (`cpu` 27.40 / `passback` 30.60 / `pcie` 23.20 tok/s, VRAM
identical across arms); the scheduler converges unaided to 0.279–0.360 on the field
shapes against the standalone bench's 0.28–0.45 by format; the two arms are
bit-exact per arm over 15 parity cases while 12 of them mix numerically different
engines; `share = 0` is byte-identical to `cpu`; and the probe's pre-fix table,
which measured the 2 ms `sync_with_deadline` poll interval instead of the kernel,
is kept because the wrong number is the kind that gets quoted.

Also records what is **not** measured: the serve/slots path, prefill, any other
host/link/arch/model, long-context decode, and the two-stream contingency.
`ShapeSnapshot` gains `min_join_ns`, and the `split:` line prints it as
`floor=0.04ms`. The floor is the quantity the controller actually compares every
join against (the smallest blocking-H2D time seen for the shape, i.e. what the copy
costs when the GPU arm finished first), so a share that looks wrong is now
diagnosable from one run without reading the state: `waited=false` with a floor at
the observed join means the share is being pushed up because no step has waited
yet, and `gpu samples 0` says the GPU rate the share rests on has never been
observed at all.
…dation

Both paths are user-facing and had no coverage — the refusal is the only thing
between a user and two engines sharing the device, and the recommendation is what
`--write` persists.

* `live_daemon_pids` is pure filesystem + `kill(pid, 0)`, so it is tested without a
  GPU or a real daemon: a live pid file is listed, a pid above the kernel's `pid_max`
  and a non-numeric file are not, `serve.pid`/`daemon.pid.bak` are not pid files, and
  a missing root is empty rather than a panic.
* The recommendation is extracted from `offload_bench_command` into
  `recommended_share`, which pins the tie behavior: one measured format is named
  outright, and equal-size measurements (the normal case — every format is measured
  over the same `buffer_mb`) take the **median**, not whichever format `max_by_key`
  happened to return last (q8_0, whose rate profile is not the model's).

`--help` for the new subcommand is rendered and checked; `--write` reuses the
`config set` path (`load_global` → `set_cli` → `write_global_toml`), which the config
tests already cover.
…e record amendment

Adds to the dated pass-back checkpoint:

* § 8 — the offload-amount sweep the record owner asked for. 9B mq4, 4/8/12/16 of
  32 layers spilled, 2 interleaved rounds per point: pass-back beats `cpu` by
  **+8.4 % / +8.5 % / +13.2 % / +13.9 %** and `pcie` by +20 % / +30 % / +39 % /
  +44 %. The advantage grows with the spill, and grows faster against `pcie`
  (which pays the whole link read per step while `cpu` and pass-back share host
  DRAM). Plus one 27B point (`qwen3.8-27b.mq3-xt`, its pinned md5, 8 of 64 spilled):
  `cpu` 14.40 / `passback` **17.10 (+18.8 %)** / `pcie` 14.40 tok/s — a bigger model
  gains more at the same *number* of spilled layers, because each spilled
  projection is a larger vector and the CPU arm's serialized host time per step is
  longer. A second 27B round was abandoned when the `pcie` arm wedged under memory
  pressure (`free` at 0 GB), the stall the headroom record already documents.
* The chart: `data-2026-10-01-offload-passback-split/mode-sweep-vs-spilled-layers.{png,svg}`
  — (a) tok/s per engine against layers spilled with a titled legend, (b) the 27B
  point with the `cpu` baseline marked, (c) the gains on their own axis so no
  percentage label sits under a data line.
* § 9 — the `offload-bench` CLI paths, unrecorded until now: the live-daemon
  refusal (exit 1, empty stdout, the pid named), `--json`, `--write` →
  `config get` → `config reset` leaving the config byte-identical.
* § 10 — the 32k-token long-range arc's fixture and intent (its numbers land in a
  follow-up commit once the three arms finish).

Amended **in place** on the record owner's explicit instruction; the file's header
now says so and marks every changed passage `[amended]`, because
`docs/perf-checkpoints/README.md` otherwise routes corrections to a new dated file.
Corrected there: the summary bullet's VRAM range (said "11308–11308 MB"; the data
are 11308–11336 across the nine arms), and the identity table's host-load
provenance (the load is `orcaslicer_main` at ~55 % of a core, not transient
YouTube, so the host was never near the plan's `uptime ≈ 0.5` and the absolute
rates are depressed — the ratios are what carry).

`benchmarks/prompts/long_range_count.txt` is the committed arc prompt (md5
`abab011aadb5cb1d3ef86a0e646cc1a3`): a self-verifying task, so a reviewer can check
the 32k-token outputs mechanically rather than reading prose.
…gains on their own axis

The record owner's complaint was that the data eclipsed the percentage labels. Two
changes fix it by construction rather than by tuning label offsets:

* panel (a) carries no in-plot text at all. Every series' four values live in the
  legend, one column per spill amount (cpu 43.8 / 28.1 / 20.5 / 16.2, and likewise
  for passback and pcie), so no number can land on a line, on another number, or
  outside the frame. The earlier in-plot version collided at two of the four x
  positions and pushed the rightmost label past the axis.
* the pass-back's gain is its own panel (c), so no percentage ever sits under a data
  line.

Panel (b) gains the cpu-baseline dashed line and a "+18.8 % over cpu" caption, and
loses a redundant legend (its x tick labels already name the engines) and a caption
that rendered red-on-red across the cpu bar's own top edge.

The chart tool that produced the figure from the raw bench JSON is committed beside
the artifact, so it is reproducible rather than a hand-drawn one-off.

No further plot work: the figure at
`docs/perf-checkpoints/data-2026-10-01-offload-passback-split/mode-sweep-vs-spilled-layers.{png,svg}`
is final.
…on, and sharpen the 27B claim

Three corrections to the amendment that landed in e728392, plus one analytic
sharpening:

* Bullet 6 quoted "19.6–28.5 GB/s" for the GPU arm and "29–73" for the CPU — the
  19.6 and the 29 are from the **pre-sync-fix** probe table that § 3 keeps only as a
  discarded comparison, so the bullet contradicted § 2's own shipped table. It now
  quotes § 2 (22.8–28.2 and 34.7–68.4 GB/s) and says what moved.
* The § 4 reading said the probe "seeds ~0.23–0.36"; 0.233 is the pre-fix seed the
  dead `sync_with_deadline` produced. The shipped seed is ~0.36, and the text now
  says so.
* § 10 pointed at a § 10.1 that does not exist yet (the arc outlives this commit).
  It now points at the briefing that does exist, with § 10.1 promised on completion.
* The 27B claim now compares at equal **spill fraction**: 8 of 64 spilled is 12.5 %
  of the model, the same fraction as the 9B's 4-of-32 point (+8.4 %), so +18.8 %
  isolates model size from spill amount. Comparing at equal spill *count* (8 of 64 vs
  8 of 32) confounds the two, and the 27B wins there too despite spilling half the
  fraction.
…attempt taught

The record owner asked for a long-horizon *text* run and then dropped that line of
work, so the arc (its prompt, its outputs and its scratch directory) is gone rather
than half-reported. § 10 now says exactly that, and replaces the promised numbers
with the two pitfalls the attempt established — both are facts a future long-budget
run needs:

* `hipfire run` has no reasoning flag. With `reasoning.mode = on` (the config
  default; the 9B is a reasoning model) two arms spent ~20 minutes of real decode
  each and wrote **1 byte**: a bare `println!()` around an empty visible answer,
  because the whole budget went into the invisible think block. `reasoning.mode =
  off` streamed text immediately (27 tok/s, matching the sweep's `cpu` point). AGENTS
  section 7 documents the `bench` side of this; `run` exposes only the config key.
* `-n` does not set an arc's length — the model's own stop does. Told to "keep going
  as long as you are allowed to", the 9B EOSed at 1000 (997 integers, 3887 bytes,
  ~3.9k tokens, 143 s), so a long run has to get its length from the prompt.

Also removes the now-unused `benchmarks/prompts/long_range_count.txt` (its md5 was
recorded only as that arc's fixture), and restores `reasoning.mode` to `on` in the
global config, which the attempt had flipped.

Closing gate: `cargo build --release` clean; `hipfire-config` 101, `hipfire-dispatch`
311 (+1 ignored), `hipfire-tui` 165, `hipfire-cli` 306 tests pass;
`scripts/check-lifecycle.py` and `scripts/check-env-docs.py` exit 0.
The passback share controller estimated the GPU arm's rate from an EWMA of
only the steps where the GPU overran the CPU multiply — a sample conditioned
on the GPU's slow tail — against a running-minimum join floor, with the
blocking D2H folded into the CPU arm's rate, and then froze permanently once
it had applied 64 adjustments.

Replace the estimator, keeping the setpoint: the DLT two-processor balance
point r_gpu/(r_gpu+r_cpu) is the correct optimum and is not the defect
(Cheng & Robertazzi's optimality principle, via arXiv:1902.01898 §II-A).

* the CPU arm's time-per-byte is sampled from the multiply alone;
* the GPU arm's is a censored (Tobit) EM estimate over both exact (overrun)
  and censored (finished-first) observations, anchored by the one-time
  probe's directly measured rate — which is also the estimator's
  identifiability anchor when every online observation is censored;
* the permanent freeze and the blind ratchet are removed, so the share keeps
  adapting for the whole decode rather than ceasing to move at 64 adjustments;
* if the automatic probe fails the share escalates upward while unanchored
  instead of sticking at the default, and the trace marks `anchor=none`;
* the split trace now prints the controller's own est_cpu/est_gpu (from the
  same state the share was computed from) rather than a per-shape mean.

Known limits, documented at the estimator and measured on the reference
fixture: the anchor is a *solo* measurement while the optimum is contended,
so the fixed point sits near 0.46 against a measured optimum of ~0.42; and
the GPU level does not re-base over a long horizon — fading the anchor lets
floor-noise overruns run the estimate away downward (est_gpu 2-9 GB/s), so
the floor must be made robust before the anchor can decay. Sources are cited
above the scheduler.
The estimator's own lag is the investigation doc's §1.2 weakness 1 ("lag, not
offset"): the GPU arm's estimate re-ran on a 64-step batch, so the share trailed
the drifting balance point by the batch length. Keep the same statistics over a
64-observation rolling window, re-estimated every 8 steps — the doc's §5 Phase 1
lever ("reduce the target EWMA's lag") applied to the estimator.

Also rewrites the module doc to the co-inference framing: pass-back is the two
offload paths (the GPU reading host-mapped weights over the link, the CPU reading
host RAM) run concurrently on one spilled step, not a fallback, and the schedule
is per step, not per layer.
Pass-back is a mixture of the two offload paths — `pcie` (the GPU reading the
host-mapped weights over the link) and `cpu` (the CPU reading host RAM) — run
concurrently on the same spilled step and divided by output rows; the share is
that schedule. Corrects every authored description: the design doc §6.2.2 (and
its heading anchor), the config schema (which feeds the env-vars table),
AGENTS.md, the CHANGELOG entry, and the benchmark-handoff note.

A step placed wholly on one engine is the schedule's degenerate point
(`share -> 0`), not a validation failure and not a fallback to a worse path.
Measured on the 9B fixture (qwen3.5:9b, budget 24 = 8/32 spilled, the pinned
gpu_offload_probe prompt, five interleaved fresh-process pairs, daemon binaries
built from 8ee62f5 and from 308f8a9): the 8-step rolling window gave 31.60
tok/s median against 32.50 for the 64-step batch, losing all five pairs, and
pushed the settled share to 0.495 against 0.464 (further from the optimum). The
estimator lag was low-pass-filtering the anchor bias, not costing throughput, so
the investigation doc's §1.2 "reduce the target EWMA's lag" lever does not hold
on this fixture.

Restores ESTIMATE_EVERY = ESTIMATE_WINDOW, at which the rolling buffer degenerates
to the disjoint batch 8ee62f5 shipped. Keeps the VecDeque mechanism and the
window constant, with the measurement recorded at the constant so the lever is not
re-attempted blind.
`weight_gemv_swiglu_residual` fuses the SiLU into the down GEMV's launch, so the
passback seam cannot see it as a `Step`: it self-split at the op level (SiLU on
the GPU, the whole GEMV+residual on the CPU) and, because its guard is the
mode-agnostic `host_mapped_cpu_capable`, that happened under passback too — so the
dense FFN down-projection, the largest non-split shape in the trace, ran wholly
on one engine (~24% of the spilled bytes on the CPU engine alone).

Under passback it now builds a `Step::GemvResidual` over the real AWQ sidecar and
hands it to `offload_split::run_with`, falling back to the whole-CPU GEMV when the
seam declines. cpu mode, pcie and resident loads are untouched: the branch is gated
on `offload_split::enabled()`, and the certified GPU fused arm is not reached.

Measured on the 9B (qwen3.5:9b, budget 24/16 = 8/16 of 32 spilled, the pinned
gpu_offload_probe prompt, interleaved fresh-process pairs, passback mode):
**+4.0 % at 8 spilled and +5.0 % at 16 spilled** against the build without it, so
the passback-vs-cpu gain at 16 spilled widens from +12.9 % to +17.6 %.

Validation: hipfire-dispatch 313 lib tests, hipfire-runtime 956; serve_harness
battery + chain on the passback arm 5/5 turns with runaway/empty/attractor 0.
…ipped one

Interleaved passback A/B, fresh process per run, daemons built from 534aff7 and
8ee62f5 (both predating the FFN-down change, so this isolates the estimator):
the shipped estimator gave 32.40 vs 31.80 tok/s at 8 of 32 spilled (-1.9%) and
19.00 vs 18.80 at 16 of 32 (-1.1%), losing 3/3 pairs at both points. The premise
of 8ee62f5 — that the shipped estimator's tail-conditioned bias drives the share
below the balance point — is smaller on this fixture than the solo-probe anchor
bias the replacement introduced (placement ~0.46 vs the shipped 0.40-0.44), so it
traded a theoretical bias for a measured one.

Restores offload_split.rs and cpu_exec.rs to 534aff7's estimator and trace and
re-applies only the co-inference framing to the module doc (that change is
orthogonal). The FFN-down co-inference from fc5c072 is untouched and still
measures +3.8% (8 of 32 spilled) / +2.6% (16 of 32) on the shipped estimator, so
the stack is now net-positive with no measured regression left in it.

Validation: hipfire-dispatch 311 lib tests; serve_harness battery + chain on the
passback arm 5/5 turns, runaway/empty/attractor/retrieval_miss 0.
The +4.0%/+5.0% figures were measured on the censored estimator, since reverted in
8f656c4. On the shipped estimator the FFN-down co-inference measures +3.8% (8 of
32 spilled) and +2.6% (16 of 32), interleaved fresh-process pairs, passback mode.
Corrects CHANGELOG.md and the design doc § 6.2.1.
… moves

The freeze gate latched a converged shape permanently — the investigation doc's
§1.2 weakness 3: it never reopens, so a long run whose context or host load moves
the optimum stays stuck at its first plateau. Removing the latch outright was tried
first and measured as a large regression (31.60 vs 32.50 tok/s at 8 of 32 spilled,
-7.0%, and -11.8% at 16 of 32, 4/4 pairs): unfrozen, the shipped estimator's
proposal wanders. So the latch stays and gains a re-open.

A latched shape keeps *measuring* — the rate EWMAs never stop — and every
APPLY_EVERY samples `observe` recomputes the proposal; if it differs from the
latched share by more than `REOPEN_EPS` (0.02) the shape unlatches, re-tracks and
can re-latch. The band sits above the largest proposal a just-latched shape can be
holding (`FREEZE_EPS / ALPHA ≈ 0.0167`), so the bands cannot overlap and the latch
cannot thrash.

Measured: latch-with-reopen vs plain latch is a tie on the static fixture (+0.6% /
+1.5% at 8/16 of 32 spilled, -0.3% at 512 tokens, all inside ±1.5%), which is the
intent — it behaves exactly as the latch until the balance point actually moves. A
`REOPEN_EPS` spread (0.005/0.01/0.02/0.05/no-reopen) confirmed the structural floor:
0.005 — below `FREEZE_EPS / ALPHA` — was the worst (-1.8% / -5.2%), while
0.01/0.02/0.05 sat inside the bench's noise. Calibrating the band from the measured
proposal scatter at runtime is the next step; a static bench cannot resolve it.

Validation: hipfire-dispatch 7 scheduler tests, including a latch-then-reopen unit
test; serve_harness battery + chain on the passback arm 5/5 turns, counters 0.
The research record behind the passback scheduling work: the perf gap (pcie
23.6 / cpu 28.1 / passback 31.1 tok/s at 8 of 32 spilled), the algorithm
survey, the constraints (no deliberate excitation, O(1)/step, cross-machine),
and the Phase 0-3 plan. Untracked until now.
The freeze latch re-opened on a single out-of-band proposal compared against a
fixed 0.02 band, both measured on the ALPHA-filtered proposal. Two things break
there. A slow host-load drift moves the balance point a little per adjustment, so
no single proposal ever leaves a band that size — on the balance-point axis
`0.02 / ALPHA ~= 0.067` — and a latch that re-latches after one step change then
sits up to 6.7 points off the drifting optimum. And one proposal past the band
can come from a host transient or a single badly sampled join, so unlatching on
it hands the controller the noise the latch exists to reject.

* `REOPEN_EPS` 0.02 -> 0.01: halves the balance-point dead zone (0.067 -> 0.033)
  while keeping a 2x margin over the freeze band (`FREEZE_EPS` = 0.005, the
  largest proposal a just-latched shape can be holding). A band at the freeze
  band re-opens on the latch's own convergence noise: the model re-opens a *clean*
  static plant nine times in 1024 steps at 0.002.
* `REOPEN_CONFIRM` = 3: unlatch only after three consecutive same-direction
  out-of-band proposals. A genuine move drives the balance point one way for many
  adjustments; noise flips sign, so the run separates them without a wider band.
  It is insurance with a cost — ~1-2 us of extra reopen latency per genuine move
  in the model, no measured throughput change — and it removes the reopens a
  noisy static host produced without it.
* `ShapeState` / `ShapeSnapshot` gain a `reopens` counter, printed on the
  `split:` trace line. `frozen` alone cannot show whether the path was exercised.

No runtime calibration of the band: the only per-shape scatter the loop holds is
the converged proposal magnitude, which the freeze guarantees is under
`FREEZE_EPS`, so a scatter-derived band would land with no margin.

Validation. The band is not field-identifiable: a four-arm A/B (0.005 without the
streak, 0.005/0.01/0.02 with it, 3 interleaved fresh-process rounds at 8 and 16 of
32 spilled on qwen3.5-9b.mq4, prompt md5 5835c71e471849b4a72e1dc8e39695e7, model
md5 296092bf1e6a45d78c1acf815eb93366) left every arm inside its own run-to-run
spread (+-6% at 16 of 32) and flipped the ordering between budgets; the 0.005
penalty an earlier spread measured did not reproduce on this host. So the value
rests on the freeze-band margin and a deterministic host-load model in
`offload_split`'s tests (step / ramp / noisy-static / transient scenarios), not on
a field win. On live serving the trace's `reopens` shows shapes unlatching 0-5
times, so the path is real.

* hipfire-dispatch 316 lib tests (11 scheduler, 4 new); hipfire-runtime offload
  and hipfire-cli `recommended_share` / `live_daemon` pass.
* serve_harness battery + chain on the passback arm: 5/5 turns, runaway/empty/
  retrieval 0. The sampled and greedy runs flag a token attractor on the prose
  turn; it reproduces on the `cpu`-mode control under identical settings (worse
  there: 3 vs 2 under greedy), so it is this model's greedy think-channel
  repetition, not the scheduler. Binaries: daemon b35c25f032f80882d757512d0f8ee98a,
  hipfire ee24d2671dc3480ac8949a43df71b156.

Not touched: the estimator's open defects (monotone `min_join_ns` floor; `r_gpu`
sampled only on a waited join) — documented in
docs/plans/partial-gpu-offload-design.md 6.2.2 "Known estimator limitations".
Phase 0 of the offload pass-back investigation: the measurement
docs/investigations/2026-10-01-offload-passback-perf-gap.md 7 names, and the gate
its 5 Phase 0 put on further scheduler tuning.

The perf-gap record's open question was whether the refused (never-split) steps
carry a disproportionate share of CPU-executed wall — the "32% of steps never
split" reading, tagged [INFERENCE] and explicitly "an assumption, not a
measurement". Bucketing the HIPFIRE_CPU_EXEC_TRACE lines by wall time on the
pinned 9B fixture answers it:

  spilled   decode tok/s   coverage (refused wall / cpu-exec wall)
   8 / 32      33.0         3.73%
  16 / 32      19.5         3.72%
  24 / 32      13.9         6.30%

The refused set is a single shape at every point: m=32 k=4096 (the DeltaNet
beta/alpha rows, ~68 KiB, under MIN_SPLIT_BYTES), 8 calls per token. At 8 spilled
that is ~2.7% of token wall. So the coverage term collapses toward zero, not
toward "most of the shortfall": the 9B's remaining gap is per-step cost on the
steps that do split (blocking copies, same-stream ordering), not coverage and not
the scheduler. The scheduler tuning itself is ce8f3e5.

Fixture: qwen3.5-9b.mq4 md5 296092bf1e6a45d78c1acf815eb93366, prompt md5
5835c71e471849b4a72e1dc8e39695e7, gfx1201, passback share auto, spec off, greedy,
128 tokens, one fresh process per point. Daemon b35c25f032f80882d757512d0f8ee98a,
hipfire ee24d2671dc3480ac8949a43df71b156. One run per point, host load
uncontrolled; the numbers are within-run wall fractions, which are load-robust,
but a second session would be needed to quote them across sessions.

New: docs/perf-checkpoints/2026-10-02-offload-passback-coverage-phase0.md.
Design doc 6.2.2 "Known estimator limitations" gains the pointer and the note
that the SHARE_MAX ratchet is latent rather than a live bug (it self-heals when a
wait reappears).
The four-arm A/B behind ce8f3e5 is a tie: every arm sits inside its own
run-to-run spread (+-6% at 16 of 32 spilled) and the ordering flips between
budgets, so the static fixture does not resolve REOPEN_EPS. The earlier 0.005
penalty (9c72282) did not reproduce. Recorded so the null is not re-run as if
open, with the arm-order caveat stated.

Arms b005c1/b005c3/b01c3/b02c3, 3 interleaved fresh-process rounds at 8 and 16 of
32 spilled, qwen3.5-9b.mq4 md5 296092bf1e6a45d78c1acf815eb93366, prompt md5
5835c71e471849b4a72e1dc8e39695e7, hipfire ee24d2671dc3480ac8949a43df71b156 held
constant across arms.
Ledger rule: the originals are immutable, corrections are new dated files.

coverage-phase0-amendment-1:
- the trace prints a shape's line only at a power-of-two call count, so the
  1024/2048/4096 calls are snapshots and the '8 per token' parenthetical (and the
  per-token wall figure derived from it) is withdrawn; the 3.7/3.7/6.3% headline
  is a ratio of two same-quantized walls and stands.
- classify the one non-split shape: m=32 k=4096 = 68 KiB < MIN_SPLIT_BYTES, so it
  is size-refused on every step, not capacity-refused; no deficit conflation here.
- [DERIVED] loop-closer: for m=12288 k=4096 the run's own rates give a balance
  floor of ~0.49 ms against a measured step wall of 0.52 ms and a balance point
  ~0.39 against a settled share 0.428, so the largest step is already at its
  balanced floor and the residual is the sync points.

latch-band-ab-amendment-1:
- 'Host load 3-6' was not measured in the 2026-10-02 session; it is carried from
  the 2026-10-01 split record. Replace with 'uncontrolled and not measured'. The
  tie is unaffected.
Amendment-1 3 computed the two-engine floor with the trace's gpu>=22.2 GB/s,
which is a LOWER bound on the GPU rate, so m*rb/(r_cpu+r_gpu) is an UPPER bound
on the achievable floor: the true floor is smaller and the slack larger than the
'~5%, already at the floor' reading. State the direction and recover the rate
from the converged balance (r_gpu = share*r_cpu/(1-share) ~ 25.6 GB/s):

  r_gpu 22.2 (lower bound) -> floor 0.494 ms -> slack 5.0% (lower bound on slack)
  r_gpu 25.6 (share-implied) -> floor 0.467 ms -> slack 10.2% (bound, not measured)

So ~5-10% of the m=12288 step is scheduling slack, and d2h+join (0.11 ms, ~21%)
is pure copy/sync. Conclusion unchanged: the residual is sync, not the share.
The share-implied rate is circular for confirming placement and is labelled as a
bound.
…se the chain

Amendment-2's 'What survives' said 'at most a few percent is scheduling slack',
contradicting its own table (5.0-10.2%). Corrected sentence states the residual is
the d2h+join sync (0.11 ms ~ 21%) plus ~5-10% scheduling slack.

Consolidated reading, with the bound directions re-checked:
- coverage 3.7/3.7/6.3%, one size-refused shape -> not the deficit;
- the floor built from gpu>= is an UPPER bound (<= ~0.494 ms), so the slack is
  >= ~5%; ~10% only under the converged-share estimate (r_gpu ~ 25.6 GB/s), which
  is an estimate, not a bound;
- d2h+join = 0.11 ms ~ 21% is pure copy/sync.
Conclusion unchanged: the residual is per-step sync cost, not share placement.
Final word on this checkpoint.
…traceable

The passback headline was split across two builds: 2026-10-01 split record
+11.7%/+13.9% (built at f907988, pre-FFN-down and pre-latch) and fc5c072's
+17.6% at 16 spilled (commit message only). Neither described HEAD.

New checkpoint: cpu/passback/pcie x {8,16 of 32 spilled} x 4 interleaved
fresh-process rounds, arm order rotated by round, identity + load in a header row.

  spilled  cpu     passback  pcie    passback/cpu  pcie/cpu
  8/32     28.65   34.10     23.90   +19.0%        -16.6%
  16/32    16.65   19.60     12.95   +17.7%        -22.2%

Consistent with the FFN-down co-inference landing on top of the old figure; the
latch work is neutral per its own band A/B. Host load 3.78 -> 8.42 (1-min),
uncontrolled and rising, so the absolutes are contended-host numbers and only the
interleaved relative deltas are readable.

HEAD 7443401, daemon b35c25f032f80882d757512d0f8ee98a,
hipfire ee24d2671dc3480ac8949a43df71b156, model 296092bf1e6a45d78c1acf815eb93366,
prompt 5835c71e471849b4a72e1dc8e39695e7.

Also: changelog gains the current-HEAD A/B pointer and the `reopens` trace field.
… to pcie on a small model

Multi-model end-to-end memory.offload_exec A/B on a spill-FRACTION grid so the
models share an axis (2B 24 layers, 9B 32, 27B 64; 2B/9B at 25/50/75% spilled,
27B at 12.5/25%), 5 rounds per point for 2B/9B and 3 for 27B, arm order rotated
by round, one fresh process per run. Median decode tok/s:

  model spilled  %      cpu    passback  pcie    pb/cpu   pb/pcie
  2B    6/24    25.0   69.90   79.30   101.00   +13.4%  -21.5%
  2B   12/24    50.0   42.10   47.10    60.20   +11.9%  -21.8%
  2B   18/24    75.0   30.00   34.00    42.60   +13.3%  -20.2%
  9B    8/32    25.0   28.10   33.60    23.80   +19.6%  +41.2%
  9B   16/32    50.0   16.50   19.90    13.20   +20.6%  +50.8%
  9B   24/32    75.0   11.60   14.10     9.10   +21.6%  +54.9%
  27B   8/64    12.5   15.00   16.90    11.70   +12.7%  +44.4%
  27B  16/64    25.0    9.60   11.30     6.90   +17.7%  +63.8%

Every passback-cpu paired round positive (5/5 at each 2B/9B point, 3/3 at each
27B point).

Findings:
- 9B is the target regime: +19.6..21.6% over cpu, rising with the spill; pcie the
  loser.
- On the 2B pass-back is the WRONG mode: pcie beats it at every spill fraction
  (+42..45% over cpu vs pass-back's +12..13%). The 2B's whole-model rates put the
  balance point ~0.59, above SHARE_MAX=0.50, so the scheduler cannot collapse to
  the GPU-only route. On the 9B/27B the balance point is 0.42..0.46, under the
  cap, and does not bind.
- A bigger model does NOT give a bigger gain (25% spilled: 9B +19.6% vs 27B
  +17.7%); the gain is roughly a rate ratio, flat in model size. What model size
  changes is which route wins.

Fixture: HEAD a7e3e1c, daemon b35c25f032f80882d757512d0f8ee98a, hipfire
ee24d2671dc3480ac8949a43df71b156, prompt 5835c71e471849b4a72e1dc8e39695e7,
models 9ed6628f/296092bf/d1292b4d. Load uncontrolled (2B 7.3->5.2, 9B 2.4->7.3,
27B 4.45->6.58), so absolutes are contended and only the interleaved relative
deltas are readable.

Also: figure + raw jsonl + chart script committed under
data-2026-10-02-offload-passback-model-spread/ (the 2026-10-01 figure's source
data lived in /tmp and is not reproducible); amends the 2026-10-02 mode-A/B record
(withdraws an unsupported outlier cause, adds the paired per-round statistic);
CHANGELOG bullet updated to the cross-model result.
…l-spread amendment

New records and corrections from the 2B/9B/27B + mq3/mq4/mq6 pass-back sweeps.

quant-spread-gfx1201.md (new): same-size, different-quant (9B, 8 and 16 of 32
spilled). pass-back vs cpu: mq3 +15.2%/+19.4%, 5/5 paired positive. Quant changes
which route wins: on mq3 `pcie` moves to +3.0%/-1.8% over cpu where on mq4 it was
-15.3%/-20.0%. mq6 is UNMEASURABLE on this build and it is a real coverage bug,
not a missing kernel: the model's tensors are legacy MQ6G256; gemv_steps(MQ6G256,
WithSwiGLUResidual) plans the UNFUSED [SiluMulRotate, GemvResidual] and
dispatch_gemv_residual already wires MQ6G256 -> gemv_hfq6g256_residual
(gemv.rs:622), but weight_gemv_swiglu_residual (hipfire-runtime/src/llama.rs:1543)
takes the FUSED launcher, and dispatch_swiglu_residual (gemv.rs:637) has no
MQ6G256 arm -> catch-all -> `unsupported gemv.swiglu_residual for /`. Plan-vs-
fused routing mismatch. Exposed alongside: the key is declared/mapped/registered
with no fused arm (coverage_tests only checks key<->table, so it cannot catch
it); the `_ =>` error carries arch/quant "" so it names no dtype (why 30 runs and
a code read were needed); and the registry advertises qwen3.5:9b-mq6 while the
shipped build cannot decode it.

model-spread amendment-1 (new): corrects four things in the committed record.
(1) the SHARE_MAX claim is REFUTED by measurement — the 2B settles at share
0.268-0.471, under the cap, so the cap is not binding; why pass-back trails pcie
on the 2B is open, not the cap. (2) the 27B's usable spill ceiling on this host is
RAM-bound (24/64 spilled thrashes 28 GB), adding the 6.25% point: cpu 20.6,
passback 22.5, pcie 17.8 (n=6 across two durable blocks), +9.0%. (3) the 9.5
outlier did not reproduce (re-run 22.6/22.6/23.5) and is kept VISIBLE in the
committed data rather than excluded. (4) baseline-dependent framing: over cpu the
gain is ~flat across size and quant; against the best single-engine route it grows
with size because that route flips identity.

Nondeterminism (design doc 6.2.2, CHANGELOG, AGENTS.md pitfalls): pass-back is
NOT bit-reproducible. Measured on the 2B: cpu x2 byte-identical, pcie x2
byte-identical, passback x2 differ, HIPFIRE_OFFLOAD_PASSBACK_SHARE=0 byte-
identical to cpu. Structural: the share is scheduled per process, so the row split
and the mixed-engine numerics vary. cpu/pcie deterministic, passback not; pin the
share in any gate that diffs output. The 2026-10-01 9B byte-identity held on its
prompts, not as a general property.

Figure data made durable and self-contained: sp_27b_extra/sp_27b_60 committed,
spread-chart.py reads its own directory, both non-reproducing/outlier rows kept.

Data: qwen3.5-9b.mq3 (sha256 c379dbbc..., 4.57 GB) and .mq6 (69b0e3b2..., 7.30 GB)
pulled and sha256-verified.
Avery Drouillard added 5 commits October 2, 2026 16:07
Follow-up to b0b4141 (kept as published rather than rewritten).

- Output determinism stated three-part: the modes are not byte-interchangeable
  (cpu vs pcie differ -- different GEMV engines), within a mode at the same layer
  split the output is deterministic (cpu x2, pcie x2 byte-identical), and
  passback alone varies per run because its scheduled row split does.
- 27B 6.25% point cites the durable blocks (cpu 20.6, passback 22.5, pcie 17.8,
  n=6) with the non-reproducing 9.5 kept visible; +9.0%.
- mq6 reframed as a plan-vs-fused routing mismatch (gemv_steps prescribes the
  unfused route; the runtime takes the fused launcher; dispatch_swiglu_residual
  has no MQ6G256 arm), plus phantom coverage and the nameless-dtype error.
- AGENTS.md pitfalls row, CHANGELOG bullet, design 6.2.2 updated to match.
Adds data-2026-10-02-offload-passback-quant-spread/quant-chart.py and
quant-vs-engine.{png,svg}, rendered from the committed jsonl (9b_mq3.jsonl +
sp_9b.jsonl). Panel (a) decode by engine for mq3 vs mq4 at 8/32 and 16/32
spilled; panel (b) pass-back's gain over cpu and over pcie, which shows the two
findings directly: on mq3 pcie is ~cpu (+12%/+22% gain over it) while on mq4 pcie
is far behind (+41%/+51%); pass-back's gain over cpu is lower on mq3
(+15%/+19%) than mq4 (+20%/+21%). Linked from amendment-1; no numbers change.
passback-scatter.{png,svg} + scatter-chart.py, rendered from the committed jsonl.
Four panels: (a) paired 1:1 passback-vs-cpu per run (every dot above the line
except the one visible 27B outlier); (b) gain vs spilled BYTES -- the curves do
not collapse, so bytes is not the sole driver; (c) gain vs spill fraction --
9B-arch peaks, 2B flat, 27B rises; (d) gain vs model depth at 25% spilled -- no
monotone size trend. Data only; no record numbers change.
…e exec-mode snapshot

Two structural fixes for the partial-offload direction the maintainer set out in
the review of the base PR: arch/dispatch code stays central, and the generic step
seam owns the execution decision rather than each call site.

- hipfire-runtime/src/llama.rs (weight_gemv_swiglu_residual): the CPU-executed
  arm hand-built a Step::GemvResidual and called offload_split::run_with directly,
  bypassing the seam. Hand the step to pipeline::execute_steps instead, so the
  seam decides cpu vs pass-back. host_mapped_cpu_capable gates on the CPU-exec
  predicate, so pcie never reaches this arm; the seam's plan_step/run_step path
  reproduces the old cpu route.

- hipfire-dispatch/src/cpu_exec.rs: collapse the two cached LazyLock<bool> mode
  predicates (cpu_exec_enabled/passback_enabled) into a single cached
  HostExecMode snapshot; the predicates become thin views.

Verification (gfx1201, qwen3.5-9b.mq4, 8/32 layers spilled, greedy, thinking off):
the cpu arm's decoded output is identical between the pre-change (HEAD^) and
current builds on a 160-token generation (sha256 57a42f72...), i.e. the seam swap
preserved the cpu route. The pass-back arm reproduces the same text as cpu on the
smoke prompt, and HIPFIRE_CPU_EXEC_TRACE shows the down-projection
(gemv m=4096 k=12288 rotated residual) split by output rows under passback and
run whole-CPU under cpu.
A cold kernel cache with parallel callers launched one hipcc per caller for the same module. Those redundant concurrent compiles are where the ROCm 7.2 clang-22 frontend crashed (parser stack fault, clang::Parser::ParsePragmaLoopHint -> SIGSEGV in Lexer::LexTokenInternal, tensor_ops.hip), failing 5 rdna-compute GPU tests in no-gpu-ci. One compile per cache key now, with the waiters re-checking the published blob and reusing it; distinct keys still compile in parallel, which is the compile_batch_for_symbols contract.

Not fixed: the clang-22 ICE itself. It is intermittent and was not reproduced by hand; this removes the redundant compiles that were present in every failing run.

Verified: cargo test -p rdna-compute --lib with HIPFIRE_KERNEL_CACHE pointed at a fresh dir, three runs, 0 failed (before the fix: 5 failed with the shared cache, 1 failed cold). ./scripts/no-gpu-ci.sh end to end on a cold cache: exit 0.
@aldrouil
aldrouil requested a review from Kaden-Schutt as a code owner October 2, 2026 22:40
Avery Drouillard added 2 commits October 2, 2026 18:45
scripts/check-crate-maps.py --check is the required "Crate maps match the tree" gate and it failed on PR warpfront#809: hipfire-arch-qwen35, hipfire-cli, hipfire-config, hipfire-dispatch, hipfire-runtime, hipfire-tui and rdna-compute had drifted -- two new modules (dispatch src/offload_split.rs, runtime src/offload_calibrate.rs), the new parking_lot dependency edge in rdna-compute, and the branch changed line/API/test counts.

Regenerated with scripts/check-crate-maps.py <crate>; only the generated marker blocks changed, the hand-written prose is preserved byte for byte. --check now reports 45 maps matching the tree.
cargo clippy --workspace --all-targets fails on hipfire-xdna with "mutable borrow from immutable input(s)" at src/lib.rs:598 (clippy::mut_from_ref is deny-by-default). Pre-existing and unrelated to this branch -- the crate is untouched here and the code is identical on master -- but it is the only hard error in the advisory clippy job, so that job stays red until it is allowed.

The &self receiver is deliberate: the BO mapping is interior-mutable by contract, the Arc keeps it alive, and the caller aliasing rule is documented on as_slice/as_mut_slice. Verified: cargo clippy -p hipfire-xdna --lib now exits 0 (4 warnings, no error).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant